(V)LMs 超越表面共现泛化能力的实证研究:来自跨模态数一致性的证据
文章背景与核心概要
本研究探讨了视觉语言模型(VLMs)究竟是学习到了抽象的语法规则,还是仅仅依赖于表层统计共现。由于标准的纯文本语言模型可以依靠文本分布线索(如 is/are 或 this/these)来推断语法数,因此要隔离出真正的泛化能力非常困难。
为了克服这一挑战,研究人员采用了一种跨模态泛化(cross-modal generalization)框架,其中关于语法数的证据完全被限制在非语言(视觉)模态中。通过向 VLMs 教授成对的新名词(通过仅在学习期间更新的新嵌入来实现),并对比由视觉线索与文本线索发出数信号的条件,作者分析了模型行为、表征动力学以及因果机制。
研究结果表明,在这两种暴露条件下,模型都表现出非平凡的跨模态泛化能力,模型内部机制对语言和非语言线索的处理方式非常相似。这为像 VLMs 这样的统计学习器能够超越表层共现并展现出真正的、兼容抽象行为的能力提供了强有力的实证证据。
Metadata
- arXiv ID: 2609.00443 [cs.CL]
- Subject: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI) - Authors: Zach Studdiford, Kanishka Misra
- Submitted: 31 August 2026
- Primary Links: View PDF | HTML Version
Metadata
- arXiv ID: 2609.00443 [cs.CL]
- Subject: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI)- Authors: Zach Studdiford, Kanishka Misra
- Submitted: 31 August 2026
- Primary Links: View PDF | HTML Version
Summary
The study investigates whether Vision-Language Models (VLMs) learn abstract grammatical rules or merely rely on surface-level statistical co-occurrences. Because standard text-only language models can rely on distributional text cues (such as is/are or this/these) to deduce grammatical number, isolating true generalization is difficult.
Summary
The study investigates whether Vision-Language Models (VLMs) learn abstract grammatical rules or merely rely on surface-level statistical co-occurrences. Because standard text-only language models can rely on distributional text cues (such as is/are or this/these) to deduce grammatical number, isolating true generalization is difficult.
To overcome this, the researchers utilized a cross-modal generalization framework where the evidence for grammatical number was restricted entirely to an extra-linguistic (visual) modality. By teaching VLMs pairs of novel nouns (via new embeddings updated solely during learning) and comparing conditions where number was signaled by visual cues versus text cues, the authors analyzed behavior, representational dynamics, and causal mechanisms.
To overcome this, the researchers utilized a cross-modal generalization framework where the evidence for grammatical number was restricted entirely to an extra-linguistic (visual) modality. By teaching VLMs pairs of novel nouns (via new embeddings updated solely during learning) and comparing conditions where number was signaled by visual cues versus text cues, the authors analyzed behavior, representational dynamics, and causal mechanisms.
The findings reveal non-trivial cross-modal generalization across both exposure conditions, with internal model mechanisms processing linguistic and extra-linguistic cues similarly. This provides robust evidence that statistical learners like VLMs are capable of generalizing beyond surface-level co-occurrence and exhibiting genuine, abstraction-compatible behavior.
The findings reveal non-trivial cross-modal generalization across both exposure conditions, with internal model mechanisms processing linguistic and extra-linguistic cues similarly. This provides robust evidence that statistical learners like VLMs are capable of generalizing beyond surface-level co-occurrence and exhibiting genuine, abstraction-compatible behavior.